Journal of Molecular Evolution
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Journal of Molecular Evolution's content profile, based on 22 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Mohanta, T. K.
Show abstract
Codon usage bias is a fundamental genomic characteristic that prefers non-random preferential use of synonymous codons. It is a major determinant of translational efficiency, gene regulation, and molecular evolution. However, the evolutionary bias and functional relevance of codon usage bias across the plant lineage is poorly defined and yet to understand what are the major factors responsible for relative synonymous codon usage (RSCU) in genomes and how codon usage bias influences the gene regulation, molecular evolution genomes. A genome-wide codon usage bias study of coding DNA sequences of 262 plant genome was conducted. It encompassed more than 4.6 billion codons from > 11 million coding sequences. Relative synonymous codon usage, codon adaptation index, codon-anticodon mapping, effective number of codon (ENC)-GC3, GC1,2-GC3, parity rule 2 (PR2-bias), molecular economy, and machine learning approaches were used for the study. It was found that codon usage bias was strongly non-random and exhibited a clear phylogenetic structuring. The higher plants favoured A/T-ending, whereas early-diverging lineages were enriched in G/C-ending codons. Analysis of RSCU, codon adaptation index, and codon-anticodon pairing indicated that translational selection is mediated by tRNA availability, contributing sustainability to these molecular patterns. Machine-learning approaches identified a small subset of codons having outsized influence on genome-wide codon usage landscapes. Further studies revealed the presence of robust inverse relationships between the effective number of codons and GC content at synonymous third positions. Neutrality analysis revealed approximately 61% of variation was driven by mutational pressure, tempered by selective constraints. Phylogenetic reconstruction showed a progressive relaxation of codon bias from algae to angiosperms while maintaining a conserved molecular economy cost of ~ 30 ATP per codon across the lineages. The study revealed codon usage bias is lineage-specific evolutionary conserved trait governed by mutation, selection, and translational optimization.
Zhang, Z.; Xu, Y.
Show abstract
Language genes can be tentatively considered as a subset of cognitive genes, although they are often discussed separately. During the evolution of SNVs (single nucleotide variations) in cognition-related genes, do language genes and cognitive genes exhibit significantly different intensities of change at several key evolutionary moments--namely, the inflection points or derivative peak positions of similarity curves drawn from multi-SNV locus bases across samples? In this study, nine distance/similarity metrics (Bray-Curtis, Cosine, Pearson, Spearman, Hamming, Jaccard, Matching, Kulczynski, and Gower) were employed to analyze 413 samples from 11 taxonomic groups, targeting SNV loci in language/cognition-related genes (13,415 effective loci, approximately 400 loci per gene), with pp6 (Homo_sapiens.GRCh38) as the reference. For each method, sample similarities (defined as 1/(1+distance)) were independently sorted in ascending order to generate raw similarity scatterplots. Due to the large sample size and representativeness, the scatter density on the similarity curves was high, and no smoothing was applied. Derivative values were calculated from adjacent similarity differences to identify peaks of evolutionary rate change (top 10 peaks per method). Combined with functional annotations of 33 language/cognition-related genes, we quantified the difference scores and occurrence frequencies of the two gene categories at the peak positions. The results indicate that cognitive-related genes exhibit slightly higher occurrence frequencies in peak windows and higher average difference scores per gene than language genes. Comparative analysis of SNVs at the peak samples and their left-side windows revealed that at positions 381-382, all nine methods shared three intersecting mutation loci, involving language genes (NFXL1, SRGAP2, SRGAP2C); at positions 355-356, there was one intersecting mutation locus, involving a language gene (SRGAP2). This suggests that certain mutations in language genes may have played a distinctive role at critical junctures in the evolution of cognitive abilities.
Parija, M.; Patra, S.; Dahanukar, N.
Show abstract
Transposable elements (TE) jump from one genomic locus to another. Since increase in their copy number is a metabolic burden for the host, TE are considered as genomic parasites. Although host-TE co-existence is regarded as an evolutionary arms race, the hypothesis is not extensively tested especially using evolutionary genomics. We provide a hypothesis testing framework to understand the distribution of TE in genic regions of the host genome, variation in the regulation of TE by host, and effect of these two factors on host-TE co-evolutionary dynamics. We test our hypothesis by understanding the distributions of potentially active TEs in the genome of 78 teleost fishes, representing major families and orders within the clade. Our analysis reveals coevolutionary arms race predicted by the Red Queen dynamics.
Perez, J.; Giunta, A. A.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of tango (tgo) in the Sep. 2015 (UC Berkeley ASM127793v1/DbusGB1) Genome Assembly (GenBank Accession: GCA_001277935.1) of Drosophila busckii. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Lieser, B. C.; Lose, B.; Kiser, C. A.; Butterfield, S.; Laschober, L.; Laskowski, L. F.; Nielsen, J.; Pulford, J.; Thompson, J. S.; Rele, C. P.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of raptor in the D. grimshawi May 2011 (Agencourt dgri_caf1/DgriCAF1) Genome Assembly (GenBank Accession: GCA_000005155.1) of Drosophila grimshawi. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Backlund, A. E.; Nielsen, J.; Pulford, J.; Suriaga, J.; Pyle, J.; McDaniel, S.; Thompson, J. S.; Rele, C. P.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of raptor in the D. eugracilis Apr. 2013 (BCM-HGSC/Deug_2.0) (DeugGB2) Genome Assembly (GenBank Accession: GCA_000236325.2) of Drosophila eugracilis. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Lawson, M. E.; Perez, J.; Giunta, A. A.; Rele, C. P.; Reed, L. K.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of tango (tgo) in the May 2011 (Broad dper_caf1/DperCAF1) Genome Assembly (GenBank Accession: GCA_000005195.1) of Drosophila persimilis. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Cornet, S.; Dennis, A. B.
Show abstract
BackgroundSynonymous mutations, once considered neutral, can affect translation efficiency through mRNA folding and splicing, generating codon usage bias. This bias is often linked to genomic GC content, which also influences gene regulation. In the parasitoid wasp Lysiphlebus fabarum, GC content was previously shown to shift between developmental stages, with larvae showing higher GC than adults. Whether this phenomenon is widespread among insects remains unknown. ResultsTranscriptomic data from six insect species spanning Diptera, Hymenoptera, and Lepidoptera was used to compare GC content between expressed genes in larvae and adults. In five species, larval transcripts exhibited higher GC content than adult transcripts. Differential expression analysis revealed that stage-biased genes displayed consistent GC shifts, and orthologous gene families with representatives across species showed particularly GC-rich larval-biased genes in Hymenoptera and Diptera. At the genome scale, modeling in 317 insect species demonstrated an association between parasitic lifestyle and reduced mean GC content in Hymenoptera and Diptera, providing a possible ecological explanation for AT-rich genomes. ConclusionsOur results show that GC content is dynamic across developmental stages, independent of overall genome composition. Stage-specific GC enrichment may reflect adaptive codon usage optimizing translation during energetically demanding life-history stages such as larval development. Furthermore, the association between parasitism and reduced genomic GC highlights how ecological lifestyle might with genome content and evolution. Lastly, this work identifies candidate genes underlying stage-specific GC bias and provides new insights into the interplay between molecular evolution, development, and parasitic adaptation in insects.
Backlund, A. E.; Nielsen, J.; Pulford, J.; Cook, B.; Anderson, J.; Robert, M.; Thompson, J. S.; Rele, C. P.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of raptor in the May 2011 (Agencourt Dere_CAF1/DereCAF1) Genome Assembly (GenBank Accession: GCA_000005135.1) of Drosophila erecta. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Hussan, J. R.; Kobro-Flatmoen, A.; Ruoff, P.; Omholt, S. W.
Show abstract
The nucleoid, which houses mtDNA within the mitochondrial matrix, is a phase-separation-driven biomolecular condensate capable of carrying out a broad spectrum of complex functions, including DNA replication, transcription, and repair. Here, we show by data-driven computational modelling that the concept of a tightly regulated intranucleoid deoxynucleoside triphoshate (dNTP) pool explains the observation that the number of mtDNA base pairs per cell is conserved in human hybrid cell lines regardless of the size of the introduced mitochondrial genome. This concept is then used to address the enigmatic observation that the synthesis rate of the short DNA strand called 7S DNA, which is part of the triple-stranded displacement loop (D-loop) found in the main noncoding region of mtDNA, increases dramatically during the cell cycle. Collectively, our quantitative analyses suggest that the mammalian mtDNA replisome uses a strictly controlled intranucleoid dNTP pool based predominantly on the synthesis and degradation of 7S DNA. One potential evolutionary explanation for this mechanism is that it offers an energetic advantage by enabling greater reliance on the salvage pathway for mtDNA replication.
Mantegazza, O.; Bertolini, L.; Leoni, G.; Colaiacovo, M.; Petrillo, M.; Bonfini, L.; Savini, C.; Ceresa, M.; Zaoui, X.
Show abstract
Although DNA Large Language Models (DNA-LLMs) offer a path to decoding genetic complexity, our ability to evaluate these models is constrained by our incomplete understanding of the very same genetic syntax and functional logic that these models are trained to learn. In this study we use single nucleotide substitutions that have or have not been observed in living organisms, to evaluate how the DNA-LLM Evo 2 interprets gene sequences from two plant model organisms, Arabidopsis thaliana and Oryza sativa japonica. Using perplexity as a measure of the model's confidence, we observe that alleles containing simulated substitutions are perceived, on average, as less likely than those observed in vivo. Although the size of the effect is modest, the effect is statistically significant and robust, suggesting that Evo 2 is aligned with our current understanding of evolutionary selective constraints. This approach is designed to be model-agnostic and species-agnostic and could serve as a generic framework for evaluating the performance of DNA-LLMs.
Thon, F. M.; Wittmann, M. J.
Show abstract
1. Plants produce a great chemodiversity, which is the diversity of specialized metabolites (SMs). These SMs are produced in complex metabolic pathways and play an important role in inter-species interactions. There are numerous hypotheses about the evolutionary processes which brought about and maintain chemodiversity. Some have been partially tested in lab and field studies. However, some of their assumptions and predictions are better tested by quantitative modeling, and so far no quantitative model has investigated the role of metabolic pathways. 2. To close this gap, we developed an individual-based model for metabolic pathway evolution. It models enzymes creating metabolites with various modifications. Enzymes undergo inheritance and mutation. We used the model to compare the screening and interaction diversity hypotheses. 3. The screening hypothesis predicts promiscuous enzymes, genetic drift, the presence of many non-beneficial metabolites, and high metabolite richness. The interaction diversity hypothesis predicts specialized enzymes, selection, the almost exclusive presence of beneficial metabolites, and situation- dependent metabolite richness. We found that the patterns predicted by the screening hypothesis did not occur, while those predicted by the interaction diversity hypothesis did. 4. This provides reason to favor the interaction diversity hypothesis over the screening hypothesis when connecting empirical results to their evolutionary context
Cabanac, S.; Dunand, C.; Mathe, C.
Show abstract
Many flowering plant species have adopted an aquatic lifestyle, contrasting with their terrestrial ancestors. Adapting to an aquatic environment required numerous evolutionary changes, including gene expansion and contraction. One of the most striking contractions has been observed in the genomes of seagrasses, where the ACO and ACS genes, involved in ethylene biosynthesis, are very few in number or even completely absent. To confirm this adaptation, we identified traces of gene loss in the genomes of four seagrass species, in the form of pseudogenes. Surprisingly, no gene loss was found in the species that had completely lost the function of ethylene synthesis, likely indicating an ancient loss of these genes. Conversely, several pseudogenes were found in the species where the ACO and ACS genes are contracting, indicating a recent and potentially ongoing process. We used the same approach on Utricularia gibba, a submerged freshwater plant, and also found a reduced number of ACO and ACS genes. In contrast, two terrestrial species closely related to seagrasses and U. gibba found a higher number of ACO and ACS genes, with no definitive evidence of gene loss. These results confirm that the loss of ethylene biosynthesis function in seagrasses is indeed linked to gene loss and suggests that it is an adaptation to a submerged rather than a marine lifestyle.
Zhang, Z.; Xu, Y.
Show abstract
This study aims to quantify the genetic similarity of different species (from fish to humans) to the human reference genome (pp6, Homo sapiens.GRCh38) based on the allele presence/absence patterns of 33 language/cognition related gene SNV loci, identify key breakpoints during evolution, and evaluate the enrichment of language and cognition genes at these breakpoints. We designed a similarity calculation method relying on binary features (four columns for A/T/C/G), adopted five difference/distance measures (Sorensen, Rogers, Nei, Reynolds, and Hellinger), and converted them into similarity values (1/(1+distance)). For each method, samples were independently ranked, the first derivative of similarity was computed, and the top 12 peaks were selected as candidate breakpoints. Results show that the similarity curves from the five methods are highly consistent (correlation coefficients >0.9), with major peaks concentrated at positions 355, 363, 381, 382, 390, 400, etc., where the corresponding samples are predominantly ancient hominins and primates. Furthermore, we defined 13 peak groups (starting positions 355-401). For each peak within a group, pairwise SNV differences between the peak apex sample and its immediate left neighbor were compared, and the intersection F_INTERSECTION (shared differential loci) was obtained. For each F_INTERSECTION, we calculated the proportions of language genes and cognition genes. In addition, we computed the differential sets between adjacent groups' F_INTERSECTION to trace the gradual emergence of new loci. In F_INTERSECTION, language genes accounted for an average of 59.5%, and cognition genes for an average of 62.9%. The proportion of language genes reached a peak at position 383 (61.2%), while cognition genes peaked at position 386 (64.9%). High frequency peak samples include c25, c27, and ja2, suggesting that language cognition genes may have undergone independent intensification during Eurasian evolution. Differential analysis between adjacent F_INTERSECTION revealed a stepwise acquisition of new loci from position 355 to 401, with three bursts of newly added loci along the entire evolutionary axis. This study provides a quantitative framework based on similarity curves, offers a novel molecular perspective for understanding the evolution of language and cognitive abilities, and highlights the potential importance of East Asian archaic hominins in the evolution of language cognition genes.
van der Ploeg, R.; Shearwin-Whyatt, L.; Grutzner, F.
Show abstract
Doublesex and mab-3 related (DMRT) genes encode a family of transcription factors central to sexual development across metazoa. DMRT genes are characterised by a highly conserved DNA binding domain (DM) while flanking regions may vary between species. Gene duplication and loss has shaped the diversity of the DMRT genes with several unresolved questions about their evolution. The most well characterised and conserved DMRT gene, DMRT1, functions as a sexual regulator universally in metazoans. In chicken, DMRT1 is located on the Z chromosome and acts as a dosage dependent primary sex determination gene. In therian mammals DMRT1 is autosomal, however, two copies are required for male development. Interestingly in the basal lineage of egg-laying mammals (monotremes), DMRT1 is localised on the X specific part of one of the X chromosomes. This provided the first evidence of a sex chromosome system with homology to the avian Z chromosome and raises questions about the function and evolution of DMRT1 in egg-laying mammals. To gain insight into the evolution of mammalian DMRT genes we performed sequence and expression analysis of monotreme DMRT genes and comparative analysis with other vertebrates. In monotremes, we identified DMRT genes 1-7, and show that DMRT8 is absent, suggesting that DMRT8 evolved in therian mammals after the divergence of monotremes. Sequence and expression analysis revealed multiple monotreme specific DMRT1 isoforms with additional protein-coding exons. The independent evolution of monotreme specific changes in DMRT1 may be the first indication of functional or regulatory differences in monotreme DMRT1. Article SummaryGenes in the Doublesex and mab-3 related (DMRT) family play important roles in sexual development across animals, but a comprehensive analysis of these transcription factors is lacking in the most basal mammalian lineage of monotremes. This comparative analysis of DMRT genes in monotremes and other vertebrates shows the conservation of DMRT genes 1- 7 but found no evidence of DMRT8 in monotremes or marsupial species, suggesting that this gene evolved in eutherians after the divergence of marsupials. The discovery of several monotreme specific isoforms and novel exons of the X linked DMRT1 reveals unique evolutionary changes in monotreme DMRT1.
Benda, P.; Uelze, L.; Brown, T. F.; Winkler, S.; Myers, E. W.; Pippel, M.; Eiseb, S. J.; Howard, A.; Pieri, M.
Show abstract
We present a genome assembly from a female Cistugo seabrae (Seabras wing-gland bat; Chiroptera; Cistugidae). The genome sequence is 1.9 gigabases (Gb) in span. The majority of the assembly is scaffolded into 25 chromosomal pseudomolecules, with the XX chromosomes assembled. The assembly has a contig N50 of 51.9 Mb and a scaffold N50 of 91.4 Mb. Species taxonomyEukaryota; Metazoa; Chordata; Craniata; Vertebrata; Euteleostomi; Mammalia; Eutheria; Laurasiatheria; Chiroptera; Yangochiroptera; Vespertilionoidea; Cistugidae; Cistugo; Cistugo seabrae Thomas, 1912 (Teeling et al., 2005; Bickham et al., 2004; Lack et al., 2010).
Lawson, M. E.; Sanow, K. A.; Fratian, M.; Matura, M.; Burton, I.; Rele, C. P.; Thompson, J. S.; Tin Chi Chak, S.; O'Rourke, K. S.
Show abstract
Gene model for the ortholog of Density regulated protein (DENR) in the May 2011 (Agencourt dgri_caf1/DgriCAF1) Genome Assembly (GenBank Accession: GCA_000005155.1) of D. grimshawi. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Nguyen Huy, T.; Dong, Y.; Ly-Trong, N.; Vinh, L. S.; Minh, B. Q.
Show abstract
Model selection is a fundamental step in phylogenetic analysis that determines the best-fit model of sequence evolution for a given multiple sequence alignment. Popular model selection methods, such as ModelFinder, rely on statistical information criteria, such as the Bayesian Information Criterion (BIC) or the Akaike Information Criterion (AIC). However, these approaches are computationally expensive and the use of information criteria has been the subject of ongoing discussion. Recently, machine learning has emerged as a promising approach for phylogenetic model selection in both nucleotide and protein sequence analyses. ModelDetector is currently the only machine learning-based method for amino acid substitution model selection. However, because ModelDetector was trained on simulated data, it does not perform well on real datasets. Another limitation is that it does not support different rate heterogeneity across sites (RHAS) models. To overcome these limitations, we introduce ProtFinder, an efficient machine learning framework for protein model selection that predicts amino acid substitution models, RHAS models, and amino acid frequency models. To enable ProtFinder to work with real datasets, we employed a transfer learning strategy consisting of three stages: (1) initial training on large-scale simulated data, (2) joint training on both simulated and real data, and (3) final fine-tuning using real data only. Experimental results show that ProtFinder outperformed ModelDetector in amino acid substitution model selection. ProtFinder achieved comparable accuracy to the maximum likelihood method ModelFinder for substitution model selection on medium and large MSAs. It performs slightly better than ModelFinder in RHAS model selection and substantially outperforms it in amino acid frequency model determination. Notably, ProtFinder is up to 1,400 times faster than ModelFinder in terms of inference time, making it particularly suitable for medium and large datasets.
Lawson, M. E.; Sanow, K.; Fratian, M.; Matura, M.; Scanlon, R.; Richard, M.; Nakhla, M.; Rele, C. P.; Thompson, J. S.; Findlay, G. D.; O'Rourke, K. S.
Show abstract
Gene model for the ortholog of Density regulated protein (DENR) in the Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Lawson, M. E.; Sanow, K. A.; Martinand, I.; Fratian, M.; Matura, M.; Rele, C. P.; Reed, L. K.; Thompson, J. S.; O'Rourke, K. S.
Show abstract
Gene model for the ortholog of Density regulated protein (DENR) in the Apr. 2013 (BCM-HGSC/Deug_2.0) (DeugGB2) Genome Assembly (GenBank Accession: GCA_000236325.2) of D. eugracilis. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.